Back

Behavior Research Methods

Springer Science and Business Media LLC

Preprints posted in the last 7 days, ranked by how well they match Behavior Research Methods's content profile, based on 30 papers previously published here. The average preprint has a 0.03% match score for this journal, so anything above that is already an above-average fit.

1
Linking continuous behavior to aesthetic enjoyment in a walkable virtual-reality museum tour: effects of agency and a painting-level analysis framework

Sklyar, Y.; Hendler, S.; Schonberg, T.

2026-09-01 neuroscience 10.64898/2026.08.27.747505 medRxiv
Top 0.1%
7.4%
Show abstract

Museum visits typically follow curator-defined routes that constrain how visitors shape their own experience, yet choice is widely held to heighten engagement, autonomy, and enjoyment. Virtual reality (VR) offers a setting in which to study these processes because it combines ecological immersion with precise, continuous behavioral measurement. We investigated (i) whether VR- derived behavioral signals are associated with self-reported enjoyment during a virtual museum tour, and (ii) whether the level of agency afforded to visitors influences enjoyment. Forty-eight adults completed a room-scale, life-size VR tour (8 * 4 m) of seven paintings from the Tel Aviv Museum of Art, each accompanied by a synchronized audio guide. Synchronized gaze and head- position streams were logged continuously (50 Hz) and segmented into painting-level viewing episodes using a trial-and-tile pipeline that intersects each painting's trial interval with an empirically defined spatial window in front of the canvas. Participants were randomly assigned to one of three agency conditions, Active (choice before every artwork), Semi-Active (choice for the first three), or Passive (fixed route),while the artwork sequence was held identical. Self- reported enjoyment at the tour and painting levels did not differ reliably across agency conditions. Among VR-derived measures, gaze engagement during the audio guide showed the clearest (though modest) association with painting-level liking, whereas locomotion and pacing measures were weak and inconsistent predictors. Agency nonetheless reliably modulated several gaze- and time-based viewing measures. The findings reveal a dissociation between subjective enjoyment and the micro-structure of viewing, and establish a reusable framework for full-tour, painting-level behavioral analysis in immersive settings.

2
Predicting Conscious Perception from Pupil's Aperture Size Using Machine Learning Techniques

Pandey, P.; Pethe, S. R.; Indrajeet, I.; Ray, S.

2026-08-31 neuroscience 10.64898/2026.08.26.747446 medRxiv
Top 0.3%
1.7%
Show abstract

Introduction: Decision making for selecting an object or a course of action from possible alternatives largely depends on our perceptual ability modulated by attention. When multiple stimuli appear close together in time, processing one stimulus can temporarily impair the processing of another due to temporal limitations of attention. Observers frequently fail to detect the second target (T2) presented within a few hundred milliseconds after the first target (T1) in a stream of stimuli, which is commonly known as attentional blink (AB). Existing theories attribute this perceptual lapse to T1 processing, distractor interference, or transient attentional gating; however, the computations underlying suppressive mechanism remains unresolved. We investigated whether pupil-size could reveal the underlying mechanisms of AB and predict conscious perception on a trial-by-trial basis. Methods: Pupil diameter and gaze locations were recorded using an infrared eye tracker. Machine learning techniques were used to classify trials when T2 was detected versus when it was not, after correct identification of T1, during an AB task from the pupil dynamics, which also yielded attentional episode (AE) associated with each element in the stream of visual stimuli when deconvolved. Results: Cross-validating classifiers achieved near-perfect accuracy not only in distinguishing but also predicting perceptual outcomes on a single-trial basis. AEs exhibited greater power when T2 was detected than when it was missed; the differential power in AEs on a logarithmic scale was highly synced with the differential pupil size. Conclusions: Collectively, these findings establish a framework for predicting attention-driven perceptual outcomes from pupil-dynamics at finer time-scale.

3
Spectral and melanopic dose calibration of consumer see-through extended-reality glasses for controlled retinal photostimulation

Gaidica, M.; Rosengart, M.

2026-08-31 ophthalmology 10.64898/2026.08.26.26361398 medRxiv
Top 0.5%
1.0%
Show abstract

Light reaching the retina is a primary regulator of human circadian physiology, acting largely through melanopsin-expressing retinal ganglion cells with peak short-wavelength sensitivity. Delivering known, repeatable retinal doses outside the laboratory is difficult because conventional light sources leave viewing geometry, gaze, and ambient conditions uncontrolled. Consumer extended-reality (XR) glasses fix a bright binocular display in constant geometry relative to the eye, but their suitability as calibrated photic stimulators has not been established. Here we validate a commercial micro-OLED XR display (VITURE Luma Ultra) for controlled retinal photostimulation. A purpose-built host application renders exact 8-bit RGB stimuli while independently controlling hardware brightness and logging all intensity-determining state; spectral radiance was measured at the retinal position of a 3D-printed phantom head with an open-source miniature spectroradiometer, anchored to absolute units by a luminance transfer calibration. The blue primary peaks at 461 nm (FWHM 43 nm), is spectrally invariant across a >10-fold intensity range, and at maximum output delivers an estimated 299 lx melanopic equivalent daylight illuminance, above consensus daytime recommendations, while remaining roughly two orders of magnitude below photobiological safety limits. The red primary is visually effective with minimal melanopic drive (melanopic DER 0.10), enabling spectrally shifted evening stimulation. Unlike the immersive virtual-reality headsets previously used for calibrated light delivery, the see-through form factor preserves the wearer's view of the surroundings--relevant for clinical monitoring in supervised settings such as the intensive care unit. These results show that consumer XR glasses can serve as a dose-calibrated platform for wearable photostimulation using an open-source measurement chain, and provide groundwork for application-layer dose-response studies.

4
Evaluating GPT-4o Model Proficiency and Clinical Reasoning for Antimicrobial Stewardship in Dentistry

Dick, M.; Madathil, S.; Patel, A.; Kapoor, H. S.; Sharma, M.; D'Souza, Z.; Hameed, S.; Abu-Samak, M.; Najirad, A.; Dwairi, D.; Radaideh, O.; Nicolau, B.

2026-09-03 dentistry and oral medicine 10.64898/2026.09.01.26361980 medRxiv
Top 0.5%
0.9%
Show abstract

Objectives: Dentists prescribe approximately one in ten antibiotics worldwide, yet antimicrobial stewardship (AMS) remains underemphasized in dental education. Large language models (LLMs) may support AMS training, but their proficiency and clinical reasoning in this context remain unclear. We evaluated GPT-4o's accuracy and clinical reasoning on dental antibiotic prescribing questions, stratified by question difficulty. Methods: We assembled 125 multiple-choice questions on dental antibiotic prescribing from eight peer-reviewed studies (2017-2023). GPT-4o answered each question and generated a clinical justification. Accuracy was assessed against source-study answer keys and examined across difficulty quartiles. Justifications were evaluated using an adapted 12-axis human-evaluation framework assessing scientific consensus, extent and likelihood of harm, inappropriate and missing content, bias, and both correct and incorrect comprehension, retrieval, and reasoning. Prophylaxis-specific questions were analysed separately. Results: GPT-4o correctly answered 72% of questions. Accuracy remained relatively stable across difficulty quartiles (78%, 78%, 65%, 70%). Experts rated 95.4% of justifications positively across the 12 axes. Comprehension, retrieval, and reasoning each exceeded 96.2% positive ratings. Missing content was the main weakness (7.8%), and 7.1% of justifications showed a moderate-to-severe potential for harm. Performance on prophylaxis-specific questions (98.1%) exceeded non-prophylaxis questions (93.0%). Conclusions: GPT-4o demonstrated moderate-to-high proficiency and clinically defensible reasoning in dental antibiotic prescribing questions. However, residual risks indicate that it is not suitable for unsupervised clinical use but shows potential as a supervised AMS educational tool.

5
Software Application Profile: A real-time surveillance system for monitoring heat exposure and its health impacts - presenting the Rio de Janeiro Heat Dashboard

de Araujo Morais, J. H.; Dias Ferreira, C.; Saraceni, V.; Medeiros de Oliveira Cruz, D.; Mateus Oliveira Aguilar, G.; Cruz, O. G.

2026-08-31 epidemiology 10.64898/2026.08.26.26361449 medRxiv
Top 0.5%
0.9%
Show abstract

Motivation: With the scaling frequency and intensity of extreme heat events across the globe, it is critical for public institutions to develop early detection systems and continuous monitoring of these events and their impacts. In Brazil, Rio de Janeiro was the first city to publish its heat protocol, with the Rio Heat Dashboard as a central component of this system. Implementation: The dashboard was implemented using R/Shiny and integrates climatic and health data from multiple sources. General features: The application comprises real-time heat exposure monitoring and automatic alert level classification, which is monitored daily by multiple municipal actors and supports activation of actions specified in the heat protocol. It also features a health impact module, which lists each heat event and its impact on mortality, and primary care and emergency visits. Availability: The source for full reproducibility is available through https://github.com/joaohmorais/RioHeatDashboard.

6
Detection of Frustration-related Operant Behavior in Rats via Machine Learning Methods

Wang, J.; Babu, A. S.; Nguyen, B.; Contreras, Y. M.; Shah, P.; Ramirez, I. C.; Green, T. A.

2026-09-01 animal behavior and cognition 10.64898/2026.08.26.747319 medRxiv
Top 0.5%
0.9%
Show abstract

Despite its strong link to neuropsychiatric conditions, frustration remains critically understudied in humans and animals alike. Therefore, there is an urgent need to develop tools to understand and therapeutically target frustration-related functions. Interestingly, humans and rats respond similarly during frustrative nonreward by increasing barpress durations. We previously validated barpress duration in rat operant tasks as a reliable measure of frustration-related behavior; however, it is wellknown that in addition to duration of responding, emotional states such as frustration alter other aspects of responding such as force of pressing. One-dimensional, static measures such as maximum force could miss rich information contained within operant data. Thus, the objective of this study is to apply machine learning (ML) to force/time profiles to discriminate frustration-related barpresses from non-frustration-related barpresses. Results showed an AUROC for FR1 (i.e., non-frustrated) vs. extinction (frustrated condition) for individual barpresses of 0.65 that improved to 0.84 with a chunk size of 10. The model generalized well to progressive ratio responding, a different kind of frustration procedure. We conclude that force/time profiling does provide utility beyond one dimensional measures of duration or force separately, meaning that we can indeed infer the internal state of frustration from behavior using ML techniques. Importantly, this project will also serve as proof-of-concept for applying ML to predict other internal states from barpress data.

7
Sedation Differentially Affects Distortion-Product And Stimulus-Frequency Otoacoustic Emissions In Chinchillas

Hauser, S. N.; Sivaprakasam, A. N.; Bharadwaj, H.; Heinz, M. G.

2026-09-01 physiology 10.64898/2026.08.26.746474 medRxiv
Top 0.7%
0.5%
Show abstract

Purpose: Otoacoustic emissions (OAEs) are used to assess outer hair cell (OHC) function. Clinical interpretation of OAE responses, however, is often limited to a present/absent binary since both physiological factors and measurement variability affect the measured OAE amplitude. Prior work showed elevated OAE responses in sedated compared to awake chinchillas, pointing to the potential influence of the medial olivocochlear (MOC) efferents on amplitudes, but this finding is inconsistent across species and OAE type. Here, we aimed to further investigate the effect of anesthesia on distortion- and reflection-type emissions in chinchillas using swept stimuli and more reliable calibration methods. Methods: Swept distortion-product (DP) and stimulus-frequency (SF) OAEs were measured in chinchillas with and without ketamine/xylazine sedation. Stimuli were presented using in-ear forward pressure level calibrations. DPOAE and SFOAE amplitudes and estimated Qerb from SFOAE group delays were compared across the two conditions. Results: We found that low-frequency DPOAE amplitudes were elevated when animals were sedated. The difference in SFOAE amplitudes was more variable across animals but appeared mildly reduced in sedated animals. Qerb estimates were slightly higher in sedated animals at some frequencies. The effect of sedation was not different across sexes. Conclusion: Taken together, these findings suggest that sedation impacts OAE measurements in chinchillas. MOC modulation could account for the present findings and differences across species. For diagnostic precision, OAE responses should be considered in the context of not only intrinsic OHC function but also extrinsic physiological processes that can modulate OHCs.

8
Capturing the imagination: mapping imagery ability across our multidimensional sense of touch

Lustenhouwer, R.; Dijkerman, H. C.

2026-08-31 neuroscience 10.64898/2026.08.26.747318 medRxiv
Top 0.7%
0.5%
Show abstract

Tactile imagery has attracted growing fundamental and clinical interest. Previous studies often investigated neural and functional similarities between imagined and actual touch. Several functional aspects of touch, such as differences between active and passive touch, between different haptic features during active touch or sensitivity of different body parts for passive touch, have also been explored in tactile imagery. Furthermore, considerable individual differences in the ability to engage in tactile imagery have been observed. However, several important aspects, involving different imagery components and a wide variety of touch qualities remain to be explored within a single comprehensive study. The current study therefore aims to provide a wide-ranging assessment of tactile imagery in terms of imagery processing components (vividness, maintenance, transformation), type of touch (active versus passive) and touch qualities (object properties for active touch, different tactile sensations across body sites for passive touch). We developed a comprehensive questionnaire containing 72 items to assess tactile imagery ability. 136 healthy participants were asked to imagine different touch types and rate imagery vividness and their ability to maintain and transform each sensation on 5-point Likert-scales. Active touch varied by object (plastic bottle, modeling clay, sponge) and property (temperature, weight, texture, resistance). Passive touch varied by body site (lip, shin, sole of the foot, lower back) and sensation (stroking, vibration, pinching). Overall, participants were able to perform tactile imagery: the vast majority reported at least some imagery across touch types. Individual variability was substantial: scores bridged both ends of the scale. Active tactile imagery differed significantly between objects, depending on tactile property. Object-property pairs with particularly strong imagery were bottle-temperature, bottle-weight and sponge-texture, whereas bottle-resistance elicited weaker imagery, as did temperature and weight for both sponge and clay. Passive tactile imagery was significantly stronger for body sites with higher receptor density (i.e. lip and foot). Imagery of stroking was significantly weaker than vibration and pinching. Active and passive imagery showed a strong, positive correlation, though some participants had relatively strong active imagery, but weaker passive imagery, or vice versa. Our findings confirm that tactile imagery ability varies across individuals and touch types, underlining the importance of a comprehensive imagery ability assessment tool specific to the tactile domain.

9
When medical credentials conflict with stated accuracy: A factorial study of source credibility and answer revision in medical LLM interactions

Wojcik, S.; Rulkiewicz, A.; Domienik-Karłowicz, J.

2026-09-01 health informatics 10.64898/2026.08.28.26361634 medRxiv
Top 0.7%
0.5%
Show abstract

Large language models perform well on medical examinations, but users routinely challenge their answers and invoke professional roles, and it is unclear what a system does when a medical credential and a stated task-specific accuracy point in opposite directions. In a factorial experiment on 480 items from four Polish specialty examination sets and three consumer large language model systems (ChatGPT, Claude, Gemini), each item and system received eleven independent conversations. Conditions crossed attributed source role (medical student, experienced specialist), stated prior accuracy on similar questions (2/10, 8/10) and suggestion correctness. The primary outcome was adoption of a prespecified incorrect option when the baseline answer matched the official key, comparing a specialist described as 2/10 with a student described as 8/10. Baseline agreement with the key was 87.2% across 15,683 analyzable conversations. The incorrect option was adopted more often from the specialist described as 2/10 than from the student described as 8/10 (10.2% vs. 7.6%; adjusted risk difference +2.82 percentage points, 95% CI +0.65 to +4.99). Estimates varied across the three systems and only one system-specific interval excluded zero. In a prespecified exploratory analysis with a shared eligibility rule, correct suggestions were adopted far more often than incorrect ones (risk difference +35.7 percentage points, 95% CI +30.8 to +40.7), indicating selective rather than indiscriminate compliance. An incorrect suggestion from a specialist with low stated accuracy was therefore slightly more influential than the same suggestion from a student with high stated accuracy, although the difference was modest and varied across systems. Agreement reached only after a user has disclosed a preferred answer should not automatically be treated as an independent second opinion, and medical large language model systems should be evaluated on how they revise answers after such disclosure, not solely on initial accuracy.

10
An interpretable, formally verified point-of-care ultrasound risk equation for difficult videolaryngoscopy: development and internal validation

Oyarzun-Silva, R. A.; Hernandez-Hernandez, P.; Fernandez-Vaquero, M. A.; De Luis-Cabezon, N.

2026-09-02 anesthesia 10.64898/2026.08.28.26361621 medRxiv
Top 0.8%
0.5%
Show abstract

Background. Videolaryngoscopy still requires adjuncts or hyperangulated rescue in a clinically important minority, and bedside screening discriminates modestly. Point-of-care ultrasound (POCUS) of the anterior airway is a promising alternative, but existing prediction models are opaque or assume a pre-specified functional form. We developed and internally validated a parsimonious, fully disclosed POCUS risk equation whose form is recovered from data and whose structural properties are machine-checked by formal proof - to our knowledge the first formally verified clinical risk predictor - following TRIPOD+AI 2024. Methods. In a prospective single-centre, single-operator cohort of 259 adults undergoing elective videolaryngoscopy (no-Easy airway 68/259, 26.3%), Sequentially Thresholded Least Squares with bootstrap stability selection (B=300) screened a 71-term library of nine POCUS features and retained a seven-term logistic equation; a two-term bootstrap-stable model was pre-specified as robustness analysis. Internal validation used 5x10 repeated cross-validation plus temporal and device hold-outs, with pre-specified overfitting and optimism assessments. Five behavioural properties of the deployed equation were machine-checked in Lean 4. Results. Two interactions met the |c|/sigma_c>2 stability criterion: skin-to-epiglottis x skin-to-hyoid-bone distance and tongue volume x sagittal tongue area. The seven-term equation reached a 5x10 cross-validated C-statistic of 0.966 (optimism-corrected 0.968) and held across temporal and device hold-outs (0.94-0.97). Calibration-in-the-large matched prevalence, with cross-validated slope 0.90 attenuating to 0.625 out-of-time; standard recalibration restored 0.92 without loss of discrimination. The pre-specified two-term robustness model reproduced this performance (C-statistic 0.964-0.968; events-per-parameter 34; shrinkage 0.99), confirming the result is not an artefact of the screening stage. Net benefit over a clinical baseline was positive across 10-50% thresholds. All five Lean 4 theorems compiled without sorry. Conclusions. A sparse, formally verified POCUS equation predicts difficult videolaryngoscopy with high internally validated discrimination and quantified, modest overfitting. Because the equation was developed in a single-operator cohort and its inputs are operator-dependent, external validation requires prior harmonisation of the measurement protocol and operator credentialing.

11
AI Video Analysis of Psychomotor Performance in EMS Education: Agreement With Human Evaluators Across Three Skills

Otte, J. H.; Cartagena, A.

2026-08-31 medical education 10.64898/2026.08.26.26361437 medRxiv
Top 1.0%
0.3%
Show abstract

Background. A primary constraint on the capacity of EMS programs to meet industry demand is psychomotor instruction and verification, requiring direct observation of each student by a qualified evaluator. Whether AI video analysis can relieve it is untested; none has been applied to EMS skill examination or compared with human examiners. Objective. To quantify human EMS evaluator inter-rater reliability and evaluate an AI video-analysis platform against it. Methods. In a prospective, fully crossed study, five certified EMS evaluators and an AI platform independently scored identical video-recorded EMT performances of cervical collar application (n=15), bag-valve-mask (BVM) ventilation (n=14), and medical assessment (n=15) on dichotomous checklists with critical-failure criteria. Agreement was assessed at item, score, and decision levels using Fleiss' kappa, Krippendorff's alpha, Gwet's AC1, and ICC(2,1)/ICC(2,k). Results. Human item agreement was moderate (kappa 0.409 to 0.467), as was single-rater reliability (ICC(2,1) 0.539 to 0.694), against good panel reliability (ICC(2,k) 0.854 to 0.919). Recorded pass/fail agreement was fair (kappa 0.297 to 0.388) and critical-failure agreement near zero for two skills (kappa 0.028, 0.119). AI alignment tracked rubric observability rather than task complexity: r = 0.857 (collar, exceeding every human), -0.173 (BVM), 0.664 (medical), and it was most lenient on two skills. Conclusions. Human evaluators are an imperfect standard, especially on critical failures. The AI was a legitimate additional rater where checklist items were discrete and visually verifiable, but not where credit required judging continuous quantities such as ventilation rate, volume, or suction duration. Defensible uses are formative and archival, not summative. These results reflect an early, non-specialist configuration: a baseline, not a limit.

12
Empowering adults to manage their hearing loss: assessing the benefits of user-controlled, smartphone-connected hearing aids.

Maidment, D. W.; Habib, A.; Gomez, R.; Benton, C.; Ferguson, M. A.

2026-09-03 otolaryngology 10.64898/2026.08.30.26361775 medRxiv
Top 1%
0.3%
Show abstract

The availability of hearing aids that can connect wirelessly to smartphone technologies via Bluetooth has grown exponentially in recent years. However, there is limited evidence assessing the benefits of user-adjustability afforded by these devices. This study aimed to assess the benefits of smartphone-connected hearing aids and an accompanying application (or app) in new and existing hearing aid users. In this single-centre, prospective, observational study, 44 adult hearing aid users (14 new and 30 existing) were recruited. Participants were fitted bilaterally with smartphone-connected hearing aids that could be adjusted by the user via an app. Self-reported outcome measures were collected at fitting and after seven-weeks of using the device in everyday life. For both new and existing hearing aid users, significant improvements in social participation, hearing-related fatigue, and hearing aid benefit and satisfaction were found. For existing hearing aid users, all outcomes were significantly better for the smartphone-connected hearing aids plus app in comparison to their existing hearing aids that did not connect to a smartphone, all with moderate-to-large clinical effect sizes (d> .6). User-controllability via the app was identified as the key benefit, and most participants (68%) reported that the app met their needs 'extremely' or 'very well'. These results suggest that, when used in conjunction with an app, smartphone-connected hearing aids can improve hearing outcomes due to greater user-controllability to improve listening. Thus, smartphone-connected hearing aids have the potential to facilitate patient-centred care, empowering the individual to successfully manage their hearing loss.

13
Default-filled outcome labels in a deployed cognitive-screening programme: an operator-level audit and the construction of twenty-four language-model arms

Ji, J.; Sun, Z.; Ying, X.; Hao, J.; Fu, Z.; Shi, D.; Kong, X.; Xu, Y.; Zhang, X.; Du, X.; Zhang, Z.; Liu, X.; Lin, P.; Wang, H.

2026-09-02 health informatics 10.64898/2026.08.28.26361585 medRxiv
Top 1%
0.3%
Show abstract

Background. Routine service databases are attractive sources of training labels for clinical prediction models, but the processes that write those labels are rarely audited before the labels are used. In a deployed community cognitive-screening programme, we audited the routine cognitive-status label, built a matrix of twenty-four model arms over the same patients under a specialist reference standard, and measured what each supervision choice bought or cost. Methods. The study cohort is the 672 individuals whose cognitive status was recorded by a titled (attending-or-above) physician, that record being the reference standard; after holding out one institution entirely, a development panel of 642 individuals at 38 institutions. The routine cognitive-status label these individuals also carry was first audited at the operator level: for each data-entry account we counted diagnoses entered and the proportion recording any impairment, and tested a competing bulk-timestamp explanation. Twenty-four arms span the supervision choices such a programme faces: an incumbent 21-variable logistic regression; local language models (Qwen2.5-1.5B/3B, Qwen3-4B/8B) zero-shot, with chain-of-thought, fine-tuned on physician labels, on routine labels with and without decontamination, or on a proxy scale-band task; preference-optimised (DPO) and reinforcement-trained (GRPO) variants; a proprietary frontier model queried zero-shot; and knowledge distillation of that frontier model into the regression and into the local 4B, using 943 teacher-labelled records from the programme's unlabelled pool. All arms are scored out-of-fold under one five-fold split grouped on registry-resolved institution clusters (no cluster spans a fold); paired contrasts use a 2,000-draw cluster bootstrap. Results. 181 operator accounts (each entering at least 100 diagnoses with zero recorded impairments) account for 45,315 rows - 40.5% of the outcome column; recorded impairment falls monotonically with account volume (15.7% for 1-9 rows to 0.7% for 500-999); a bulk-timestamp explanation was tested and refuted, identifying the write-time column as a migration artefact. Under the specialist standard, no locally fine-tuned arm beat the incumbent regression (AUROC 0.926): physician-label SFT reached 0.924 (4B), DPO 0.881, and GRPO 0.789; the pre-registered two-stage proxy-then-RL recipe was worse than its single-stage contaminated baseline (-0.030, 95% CI -0.077 to -0.004). Chain-of-thought reduced discrimination at every size (-0.072, -0.080, -0.041 at 1.5B/3B/4B; -0.012, n.s., at 8B). The frontier model scored 0.932 (vs. regression +0.007, n.s.). The distilled 4B reached 0.940 - above the incumbent (+0.014, 0.004 to 0.031) and above its own teacher (+0.008, 0.001 to 0.017) - with near-teacher calibration; it reached the teacher's level by 50 teacher labels and changed little beyond 200. Conclusions. The audit and the arm matrix support one deployment recipe: audit the routine label at the operator level before training on it; do not expect fine-tuning, preference optimisation, or reinforcement learning on a few hundred specialist cases to beat a well-calibrated regression; and if a frontier model is available but undeployable, spend a bounded number of queries on it as a labelling instrument and distil. A companion paper uses these frozen predictions to quantify how evaluation design choices compare with model choice.

14
Limits of Trial-Adaptive Neural Language Fusion Across Large Language Models in P300 Brain Computer Interfaces

Gorenshtein, A.; Omar, M.; Jia, E. L.; Adiniaev, Y.; Daniel, O.; Kruskal, J.; Ahmed, M.; Brook, O. R.; Klang, E.; Barash, Y.

2026-09-03 neurology 10.64898/2026.08.30.26361777 medRxiv
Top 1%
0.2%
Show abstract

Objective: Published P300-speller fusion schemes fix prior trust regardless of trial reliability; we tested whether a reliability estimate improves on it. Methods: We reanalyzed 3,373 archived P300-speller selections from 47 people with ALS (BigP3BCI). A fair, matched-search-space comparison, tuning both a fixed weight and an adaptive policy out-of-fold, was evaluated across 22 evaluable language-model priors up to 46.7B parameters. Two representative priors, GPT-2 and a classical 5-gram, additionally received detailed naive and mechanistic analyses. Results: No prior's 95% CI favored adaptive fusion under the fair comparison, despite unexploited oracle headroom at every scale. Under GPT-2, the naive comparison was significantly worse for adaptive fusion; both anchors converged to a degenerate or near-degenerate fair-comparison solution. For the representative anchors, three further controllers failed to convert that headroom into benefit; the fixed-fused posterior's output probability outperformed the best controller for flagging errors (2.8- to 3.8-fold enrichment). Conclusion: A tuned fixed weight is a difficult-to-beat default across the tested scale range; reliability estimation gave no deployable adaptive advantage. Significance: Adaptive weighting should be validated against a fairly tuned baseline across model families and scales; in this dataset, the fused output's confidence identified high-risk selections better than the tested purpose-built ranker.

15
Returning APOE and pTau-217 Results: the eSMARTER Randomized Noninferiority Clinical Trial

Langbaum, J. B.; Erickson, C. M.; Langlois, C.; Wood, E. M.; Egleston, B. L.; Harkins, K.; Mim, R.; John, S.; Brown, C.; Brown, S.; Howe, S.; Cacioppo, C.; Eppelmann, L.; Enos, J.; Salata, H.; DeSantiago, D.; Largent, E. A.; Reiman, E. M.; Denkinger, M. N.; Ashton, N. J.; Roberts, J. S.; Karlawish, J.; Bradbury, A. R.

2026-09-01 neurology 10.64898/2026.08.27.26361535 medRxiv
Top 1%
0.2%
Show abstract

Importance: Patients are increasingly learning Alzheimers disease (AD) genetic and biomarker results through electronic health portals. Evaluation of alternative scalable delivery models for return of AD risk information is needed to best support patient understanding and psychological well-being. Objective: To determine whether a patient-centered digital platform is comparable to clinician-mediated telehealth sessions for returning APOE and plasma pTau-217 results on outcomes of knowledge and psychological well-being. Design: The Evaluation of Self-Mediated Alternatives for Risk Testing Education and Return of Results (eSMARTER) study was a noninferiority trial of a patient-centered digital platform compared to clinician-mediated disclosure of APOE genotype and optional pTau-217 disclosure. Setting: Decentralized, fully remote trial enrolled participants in the contiguous United States (U.S.) between October 2024 and February 2025, with follow-up completed in November 2025. Participants: Eligible participants were aged 60-80 and had previously undergone APOE genotyping (without disclosure) via the GeneMatch program, passed psychological screening, had internet access, and were English-speaking. Interventions: Participants were randomized, 2:1, to the eSMARTER digital platform or clinician-mediated disclosure of APOE genotype. Following the 6-month post-APOE assessment, participants were offered optional pTau-217 disclosure via the same randomized modality. Main Outcomes and Measures: Primary outcomes at 1-7 days following APOE disclosure included changes in anxiety, disease-specific distress, and AD-related knowledge within a priori non-inferiority margins. Results: 674 persons (mean [SD] age 68 [4.7] years; 451 [67%] female; mean [SD] telephone MoCA=19 [2]) were eligible and provided demographic information. 651 participants were randomized to clinician-mediated (n=216) or digital disclosure (n=435) and completed APOE disclosure (66 [10%] APOE4 homozygotes, 377 [58%] heterozygotes, 208 [32%] non-carriers). 604 participants completed the study; 500 completed optional pTau-217 disclosure. Baseline characteristics were balanced across groups. At 1-7 days following APOE disclosure, scores on AD-related knowledge, PROMIS Anxiety, and disease-specific distress measures met non-inferiority. Conclusions and Relevance: Disclosure of APOE genotype by the eSMARTER digital platform is non-inferior to clinician-mediated telehealth disclosure. No significant between group differences were found following disclosure of pTau-217 results. Together, these results suggest that this digital platform may provide an evidence-based scalable approach for returning AD genetic and biomarker results.

16
The use of computerised testing to assess cognitive performance in people with HIV in South Africa

Edmond, E. C.; Dreyer, A. J.; Winston, A.; Khoo, S. H.; Joska, J.; Nightingale, S.

2026-08-31 hiv aids 10.64898/2026.08.27.26361083 medRxiv
Top 2%
0.1%
Show abstract

Background Computerised cognitive testing may address the global challenge in identifying cognitive changes in people living with HIV scalably and affordably. We assessed a computerised battery (CB) of cognitive tests, in a prospective cohort (CONNECT) of people with HIV in a low-income peri-urban area of Cape Town, South Africa during a national programmatic switch from efavirenz- to dolutegravir-based antiretroviral therapy (ART). Methods We recruited 170 people with HIV and 91 people without HIV (controls) (140[82%] and 41[45%] followed up). The CB and gold-standard pen&paper cognitive testing (P&P) were performed at both timepoints. Technology familiarity/use questionnaire data were also collected. We compared performance in detecting lower group-level cognitive performance associated with efavirenz treatment. Furthermore, the CB was compared to P&P in classifying individuals with low cognitive performance, correlation of global test scores and domain-level scores between batteries, and practice effects between timepoints. Exploratory principal component analysis was also performed. Results People with HIV on efavirenz at baseline had lower performance on the computerised battery than controls, {Delta}T=2.6, p=0.0047. This difference was lost after switching to dolutegravir-based ART at follow-up. CB and P&P global T were moderately correlated (R2=0.203, p<0.001), and the CB performed moderately in classification of low cognitive performance against the gold standard (AUC 0.70, sensitivity 0.52, specificity 0.76, PPV 0.40, and NPV 0.84). Selecting the first three principal components improved both classification of low cognitive performance (AUC 0.77) and correlation strength with P&P global T (R2=0.3, p<0.001). The CB did not show practice effects. Most participants owned a mobile phone (95%, 85.9% of these smartphones). Performance was better in smartphone owners ({Delta}T=1.8) and computer owners (23%, {Delta}T=1.8). Conclusions Delivering computerised cognitive testing was feasible in this low-income southern African setting. The CB showed reasonable construct validity (detecting known lower cognitive performance associated with efavirenz-ART) and may detect broad cognitive characteristics such as processing speed and accuracy. However, correlation of CB results with gold standard P&P testing was low-moderate and may limit its applicability as a diagnostic tool. This might be improved by including a wider range of cognitive domains tested in the CB, or data driven analysis. Brief CBs may fulfil an initial screening role, followed by more detailed clinical assessment.

17
Can a General-Purpose Coding Agent Analyze a Production Hospital Data Warehouse?

Marshall, N. P.; Haberkorn, W.; Faulkenberry, J.; Palakkattil, G.; Nateghi Haredasht, F.; Schwenk, H. T.; Chen, J. H.; Morse, K.

2026-09-04 health informatics 10.64898/2026.09.02.26362008 medRxiv
Top 2%
0.1%
Show abstract

Background. Health systems answer most questions by having expert analysts hand-write queries against a complex electronic health record data warehouse, a slow, resource-intensive process. Whether an autonomous coding agent can do this accurately is unknown. Methods. In a single-center quality-improvement evaluation, we posed ten questions about a common pediatric infection, acute otitis media. The questions went to analysts, whose adjudicated answers were the reference, and to an autonomous coding agent (OpenAI Codex), which wrote and ran read-only queries on a full copy of the production data warehouse (Epic Caboodle). We ran the agent under four conditions: an autonomous baseline (each question answered three times at two reasoning-effort settings), a variant in which it listed its assumptions, interactive analyst feedback, and reuse of a corrected definition across related questions. Outcomes were accuracy, patient-level agreement (F1), reproducibility, and cost. Results. Working autonomously, the agent wrote valid queries and never fabricated data, but rarely produced the exact answer. At medium effort, it came within 5% of the reference on 27 of 30 runs but exact on only 10. Higher effort produced no improvement. Reproducibility was the greater weakness, with the three runs returning an identical answer on only 3 of 10 questions. When a run matched the reference, it had found the same patients (F1 = 1.00); the exception was an over-counted procedure (F1 = 0.72). With analyst feedback on four questions, it answered two exactly, and reusing the corrected definition on related questions restored reproducibility and accuracy. Conclusions. A general-purpose coding agent matched an adjudicated analyst reference on most routine questions but not reproducibly, because key definitions depended on warehouse knowledge that the data dictionary omits. Supplying that knowledge, by prompt or feedback, restored reproducibility. Under expert analyst supervision, the agent is already a capable drafting aid and a promising step toward broader hospital analytic support.

18
Higher rewards lead to more accurate flower detection and increased contrast sensitivity in the bumblebee Bombus terrestris

Robert, T.; Flett, E.; Le Lay, H.; Nicolas, M.; Nityananda, V.

2026-08-31 animal behavior and cognition 10.64898/2026.08.26.747241 medRxiv
Top 2%
0.1%
Show abstract

In vertebrates, top-down visual attention is a cognitive process where internal goals modulate the tuning of peripheral sensory systems. This leads to increased perceived contrast to both goal-relevant objects and areas of the visual field that are attended. Such a system would also be beneficial to bees, enabling them to detect and recognise the most profitable flowers in their environment. We tested whether bumblebees possess a top-down attentional system resembling that seen in vertebrates. We trained two groups of bees to collect rewards under high contrast targets. To potentially induce a difference in attention while searching for the targets, one group received a higher concentration of sucrose rewards compared to the other. During tests, the targets were presented with a series of lower contrasts to measure the contrast sensitivity curves of the bees induced by the different learnt reward levels. We predicted a stronger effect of any attention-like process on contrast sensitivity in the high reward group. We also repeated this experiment with the neonicotinoid pesticide imidacloprid dissolved in the sucrose rewards to test whether this affects bee attention. Across all test contrasts, higher rewards significantly increased bee accuracy when locating targets, lowered contrast thresholds and reduced the latency to make first choices. Imidacloprid reduced bee accuracy but did not influence first choice latency. These results suggest that learnt floral rewards can influence bee behavioural contrast sensitivity in a manner resembling vertebrate top-down attention and that imidacloprid may modulate this through effects on their nervous system.

19
Large language model-augmented implicit surgical video review

Zhang, Z.; Qadir, M. I.; Ramchand, R.; Belwadi, M.; Ball, R. P.; Konstantinopoulos, K.; Abbey, E. M.; Ernsberger, K. T.; Guzman, M. J.; Hendren, S.; Holcomb, B. K.; Robb, B. W.; Stankowski, T.; Waters, J. A.; Stefanidis, D.; Bilimoria, K. Y.; Mohanty, S.; Kolbinger, F. R.

2026-08-31 surgery 10.64898/2026.08.25.26361071 medRxiv
Top 2%
0.1%
Show abstract

Surgical video interpretation is a promising medical artificial intelligence application. However, no existing video annotation method preserves the spatiotemporal complexity of surgeon reasoning. Here we show that verbal reasoning and visual attention can be converted into structured, machine-actionable records of intraoperative behaviours. Our method decomposes transcribed verbal commentary into video-anchored semantic feedback chunks, which are classified via a large language model, with spatial grounding to surgical scenes via eyegaze or cursor tracking. We demonstrate method validity and scalability on structured and unstructured annotation tasks. For quality feedback on full-length colorectal procedures, the method reached near-human fidelity for chunking (mean cosine similarity: 0.95, SD: 0.01) and semantic classification across observations (mean Cohen's kappa: 0.71, SD: 0.07) and evaluative triggers (mean Cohen's kappa: 0.67, SD: 0.14), with excellent usability ratings. For structured critical view of safety assessment in laparoscopic cholecystectomy, implicit annotation yielded excellent agreement with explicit reviewer ratings (Cohen's kappa: 0.83, 0.49 and 0.81 across three criteria). We anticipate this method will advance surgical data science by enabling scalable construction of meaningfully annotated surgical video datasets.

20
Identification of the Minimal Clinically Important Difference (MCID) for Childhood Autism Rating Scale Second Edition (CARS2) in children with ASD

Vyshedskiy, A.; Pavoski Poloni, L. E.; Schmiedel Fucks, A.; Khokhlovich, E.; Fucks, E.; Schmiedel, A.

2026-09-04 pediatrics 10.64898/2026.08.31.26361823 medRxiv
Top 2%
0.1%
Show abstract

Purpose: In clinical trials, treatment efficacy is commonly assessed by comparing control and treatment groups. However, in large samples, even small and clinically trivial differences may achieve statistical significance. Accordingly, the Minimal Clinically Important Difference (MCID) is used as a threshold to determine whether statistically-significant effects are also clinically meaningful to patients. The objective of this study was to estimate the MCID for the Childhood Autism Rating Scale Second-Edition (CARS2) using the Patient Impression of Change (PIC) as an external anchor. Methods: Single-item PICs are not well suited to characterizing improvement in a multifaceted disorder such as ASD. Accordingly, the 77-item Autism Treatment Evaluation Checklist (ATEC) was used as a multi-item PIC. CARS2 and ATEC were administered concurrently to 62 children with ASD, aged 1.8-7.9 years, with assessments conducted six months apart. Results: The correlation between changes in CARS2 and ATEC total scores was 0.41-0.44 (p<0.0001), supporting the use of ATEC as an anchor measure. Two anchor-based methods yielded MCIDs of 2.39-2.82 CARS2 points. Two distribution-based methods produced MCIDs of 1.33-3.32. Conclusions: Taken together, these approaches suggest that a between-group difference of 2.6 CARS2 points (the midpoint of the anchor-based estimates) may serve as the MCID in children with ASD.